Papers with fine-grained understanding

8 papers
RAVEN++: Pinpointing Fine-Grained Violations in Advertisement Videos with Active Reinforcement Reasoning (2025.emnlp-industry)

Copied to clipboard

Challenge: Recent advances in large language models have improved the detection of non-compliant content, but critical gaps persist in fine-grained understanding, explainability, and generalization.
Approach: They propose a framework that combines active reinforcement learning, fine-grained violation understanding and progressive multi-stage training.
Outcome: The proposed framework outperforms general-purpose LLMs and specialized models in fine-grained violation understanding, explainability, and generalization.
Do Video Language Models really understand the video contexts? (2025.naacl-srw)

Copied to clipboard

Challenge: Recent advances in VideoQA performance have shown that visual language models are effective but the processes of understanding and reasoning in VLMs remain under-explored.
Approach: They propose a framework that incorporates a fine-grained question generation and answering process to measure how well VLMs understand video question answering tasks.
Outcome: The proposed framework incorporates a fine-grained question generation and answering process to measure how well the responses generated by VLMs align with what the model understands.
TemporalVLM: Video LLMs for Temporal Reasoning in Long Videos (2026.findings-acl)

Copied to clipboard

Challenge: Several video understanding applications require the ability of temporal reasoning.
Approach: They propose a video large language model for temporal reasoning and fine-grained understanding in long videos.
Outcome: The proposed model outperforms existing methods in time and motion studies and temporal action segmentation evaluations.
TUNA: Comprehensive Fine-grained Temporal Understanding Evaluation on Dense Dynamic Videos (2025.acl-long)

Copied to clipboard

Challenge: Existing benchmarks for video understanding often focus on specific aspects, overlooking the holistic nature of video content.
Approach: They propose a temporal-oriented benchmark for fine-grained understanding on dense dynamic videos with two complementary tasks: captioning and QA.
Outcome: The proposed model performs well on diverse video scenarios and dynamic videos, with interpretable and robust evaluation criteria.
Timeline-based Sentence Decomposition with In Context Learning for Temporal Fact Extraction (2024.acl-long)

Copied to clipboard

Challenge: Recent research on temporal fact extraction fails to establish time-to-fact correspondences in complex sentences.
Approach: They propose a timeline-based sentence decomposition strategy using large language models with in-context learning to extract temporal facts from natural language text.
Outcome: The proposed method achieves state-of-the-art on a complex temporal fact extraction dataset.
Understanding Fine-grained Distortions in Reports of Scientific Findings (2024.findings-acl)

Copied to clipboard

Challenge: a fine-grained understanding of how scientific findings are reported is crucial, says a new study . a recent study found that tweets distort scientific findings more often than news reports .
Approach: They propose to annotate 1,600 scientific findings from academic papers paired with corresponding tweets . they also establish baselines for automatically detecting these characteristics .
Outcome: The proposed method outperforms few-shot prompting in detecting distortions in unpaired data.
GHAN: Graph-Based Hierarchical Aggregation Network for Text-Video Retrieval (2022.emnlp-main)

Copied to clipboard

Challenge: Existing approaches to text-video retrieval are limited due to structural and semantic differences between text and video.
Approach: They propose an end-to-end graph-based hierarchical aggregation network for text-video retrieval according to the hierarchy possessed by text and video.
Outcome: The proposed model achieves Recall@1 of 73.0%, 65.6%, and 64.0% better than the current state-of-the-art model.
YouMakeup: A Large-Scale Domain-Specific Multimodal Dataset for Fine-Grained Semantic Comprehension (D19-1)

Copied to clipboard

Challenge: Multimodal semantic comprehension has attracted increasing research interest recently such as visual question answering and caption generation.
Approach: They propose to use a large-scale multimodal instructional video dataset to support fine-grained comprehension research in specific domain.
Outcome: The proposed dataset contains 2,800 videos from YouTube, spanning more than 420 hours in total.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations